Papers with cutting-edge Multi-Modal Large Language Model
MM-IGLU: Multi-Modal Interactive Grounded Language Understanding (2024.lrec-main)
Copied to clipboard
| Challenge: | In human-robot interaction, a robot interprets user commands related to its environment, aiming to discern whether a specific command can be executed. |
| Approach: | They propose to integrate user statements with environment's description to create a multi-modal interactive Grounded language understanding model that integrates both visual and textual data. |
| Outcome: | The proposed model integrates user’s statement with environment’s description and a cutting-edge Multi-Modal Large Language Model merges both visual and textual data. |